Papers with multilingual multi-modal encoder
Translation-Enhanced Multilingual Text-to-Image Generation (2023.acl-long)
Copied to clipboard
| Challenge: | Existing models for text-to-image generation are mostly based on the English language due to the lack of annotated image-caption data in other languages. |
| Approach: | They propose to use a multilingual multi-modal encoder to bootstrap mTTI systems that can be translated into other languages. |
| Outcome: | The proposed approach mitigates the language gap and improves on standard mTTI datasets. |